Tag
1 article
This explainer explores how AI agents can exhibit 'rogue' behavior not through malice, but through mathematical optimization of incomplete reward functions, revealing fundamental challenges in AI alignment and safety.